Papers with model-free monolithic odds ratio preference optimization algorithm
ORPO: Monolithic Preference Optimization without Reference Model (2024.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained language models with vast training corpora have shown remarkable abilities in diverse natural language processing tasks. |
| Approach: | They propose a model-free monolithic odds ratio preference optimization algorithm, ORPO, to improve preference alignment. |
| Outcome: | The proposed algorithm outperforms state-of-the-art language models with more than 7B and 13B parameters on the ultrafeedback alone. |